Papers with statistical tests

10 papers
Inspecting Soundness of AMR Similarity Metrics in terms of Equivalence and Inequivalence (2024.starsem-1)

Copied to clipboard

Challenge: Existing Abstract Meaning Representation (AMR) similarity metrics have less investigated their soundness .
Approach: They propose a new experimental method to evaluate soundness of AMR similarity metrics in terms of equivalence and inequivalentity.
Outcome: The proposed method satisfies the soundness criteria of existing AMR similarity metrics and improves them by proposing a revised metric, SMATCH .
Large-Scale Hate Speech Detection with Cross-Domain Transfer (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets for hate speech detection are limited due to the labor cost.
Approach: They construct large-scale tweet datasets for hate speech detection in English and a low-resource language, Turkish, consisting of human-labeled 100k tweets per each.
Outcome: The proposed datasets outperform conventional bag-of-words and neural models by at least 5% in English and 10% in Turkish for large-scale hate speech detection.
We Need to Talk about Standard Splits (P19-1)

Copied to clipboard

Challenge: Existing methods to evaluate systems with a held-out test set are insufficient for system comparison.
Approach: They propose to use multiple random splits to compare performance of systems . they replicate results on standard split but fail to reproduce some rankings .
Outcome: The proposed method is based on multiple random splits to replicate results with a set of part-of-speech taggers.
Matics Software Suite: New Tools for Evaluation and Data Exploration (L18-1)

Copied to clipboard

Challenge: Numerous works propose interfaces or frameworks to build, explore and visualize corpora of annotated data.
Approach: Matics proposes a dataframe data model for exploring annotated data and evaluation results.
Outcome: The tools already run on several Natural Language Processing tasks and standard annotation formats, and are under on-going development.
Towards Robust Comparisons of NLP Models: A Case Study (2025.coling-main)

Copied to clipboard

Challenge: Existing statistical tests to compare the test scores of different NLP models have been proposed to account for nuisance factors such as noise, randomness, or hyperparameter values.
Approach: They propose a regression analysis which isolates the effect of nuisance factors from the effects of the models’ capabilities.
Outcome: The proposed model is able to show that the difference between BioLinkBERT and MSR BiomedBERT is 7 times smaller than previously reported.
DHP Benchmark: Are LLMs Good NLG Evaluators? (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly serving as evaluators in Natural Language Generation (NLG) tasks.
Approach: They propose a framework that measures the discernment of Large Language Models (LLMs) across diverse NLG tasks.
Outcome: The proposed framework provides quantitative discernment scores for LLMs across four NLG tasks.
Awes, Laws, and Flaws From Today’s LLM Research (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are a powerful technology that can follow instructions and output coherent, persuasive text.
Approach: They examine the scientific methodology behind large language model (LLM) research and cross-validate it with arguments at the centre of controversy.
Outcome: The authors cross-validate 2,000 research works released between 2020 and 2024 based on criteria typical of what is considered good research and find that conference checklists are effective at curtailing some of these issues, but balancing velocity and rigour in research cannot solely rely on these.
Downstream Trade-offs of a Family of Text Watermarks (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) can generate humanlike responses to a variety of requests like writing emails, translating or summarizing content.
Approach: They evaluate the performance of large language models (LLMs) watermarked using three different strategies over a diverse suite of tasks including those cast as k-class classification (CLS), multiple choice question answering (MCQ), short-form generation (e.g., open-ended question answering) and long-form generator (eg. translation)
Outcome: The proposed models can cause significant drops in their effectiveness across a variety of tasks including CLS, MCQ, short-form generation and translation tasks.
CONTESTS: a Framework for Consistency Testing of Span Probabilities in Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Language model scores are often treated as probabilities, but their reliability as probability estimators has mainly been studied through calibration, overlooking other aspects.
Approach: They propose a framework to assess model reliability across interchangeable completion and conditioning orders by performing statistical tests on real and synthetic data to eliminate training effects.
Outcome: The proposed framework assesses the consistency of model predictions across interchangeable completion and conditioning orders on real and synthetic data to eliminate training effects.
Exploring Intra and Inter-language Consistency in Embeddings with ICA (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that ICA can reveal universal semantic axes across languages but lack verification of consistency of independent components within and across languages.
Approach: They propose to use independent component analysis to identify independent components that are more interpretable than PCA to find universal semantic axes.
Outcome: The proposed framework ensures the reliability and universality of semantic axes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations